Papers with language technology
Summarization of Dialogues and Conversations At Scale (2023.eacl-tutorials)
Copied to clipboard
| Challenge: | Conversations are the natural communication format for people. |
| Approach: | This tutorial will survey the cutting-edge methods for summarizing written and spoken conversation. |
| Outcome: | This tutorial will examine the cutting-edge methods for summarizing written and spoken conversations, covering key sub-areas whose combination is needed for a successful solution. |
Ethical Considerations for Low-resourced Machine Translation (2022.acl-srw)
Copied to clipboard
| Challenge: | a paper examines the ethical implications of machine translation for low-resourced languages . a value scenario illustrates potential harms that low-rsourced language communities may face . |
| Approach: | They propose to use Armenian as a case study to investigate ethical implications of machine translation for low-resourced languages. |
| Outcome: | The proposed model is based on a value-scenario model of machine translation for low-resourced languages . the model is used to identify potential harms that low-income speakers may face . |
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)
Copied to clipboard
Michael Rosner, Sina Ahmadi, Elena-Simona Apostol, Julia Bosque-Gil, Christian Chiarcos, Milan Dojchinovski, Katerina Gkirtzou, Jorge Gracia, Dagmar Gromann, Chaya Liebeskind, Giedrė Valūnaitė Oleškevičienė, Gilles Sérasset, Ciprian-Octavian Truică
| Challenge: | Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources. |
| Approach: | They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied . |
| Outcome: | The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem . |
How many words does it take to understand a low-resource language? (2025.naacl-srw)
Copied to clipboard
| Challenge: | We evaluated the documentation needed to create a sentence embedding space using widely spoken languages. |
| Approach: | They propose to use widely spoken languages as a proxy for low-resource languages to evaluate the documentation needed to create a sentence embedding space. |
| Outcome: | The proposed language model can be used to improve the performance of sentences embedded in low-resource languages. |
Assessing Monotonicity Reasoning in Dutch through Natural Language Inference (2023.findings-eacl)
Copied to clipboard
| Challenge: | a novel dataset for natural language inference (NLI) is used to study monotonicity reasoning in Dutch. |
| Approach: | They investigate monotonicity reasoning in Dutch using a novel dataset . they find that models struggle with downward entailing contexts . |
| Outcome: | The proposed dataset shows that models struggle with downward entailing contexts, and argue that this is due to a poor understanding of negation. |
Building and curating conversational corpora for diversity-aware language science and technology (2022.lrec-1)
Copied to clipboard
| Challenge: | Language resources that capture language use in its natural habitat of social interaction are rare despite the obvious merits of studying the very environment where we all learn and use it everyday. |
| Approach: | They propose to build an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. |
| Outcome: | The proposed pipeline can be used to collect and curate conversational corpora in 67 languages and varieties from 28 phyla. |
Machine Translationese: Effects of Algorithmic Bias on Linguistic Complexity in Machine Translation (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing studies have shown that existing models amplify biases observed in training data. |
| Approach: | They propose to use MT and NLP to amplify biases observed in training data to investigate how bias amplification might affect language in a broader sense. |
| Outcome: | The proposed model amplifys biases observed in training data and could lead to an artificially impoverished language, the authors show. |
Towards Zero-shot Language Modeling (D19-1)
Copied to clipboard
| Challenge: | a number of natural questions have been asked about the inductive biases of neural networks on core NLP tasks. |
| Approach: | They construct an informative prior for held-out languages on a task of character-level, open-vocabulary language modelling. |
| Outcome: | The proposed model outperforms baseline models with an uninformative prior in both zero-shot and few-shot settings, showing that it is imbued with universal linguistic knowledge. |
Mean Machine Translations: On Gender Bias in Icelandic Machine Translations (2022.lrec-1)
Copied to clipboard
| Challenge: | a study conducted on Icelandic translations in the translation systems Google Translate and Véling.is . results show a pattern which corresponds to certain societal ideas about gender. |
| Approach: | They examine how gender bias appears in English-Icelandic translations . they conducted a study on Icelandic translation in the translation systems Google Translate and Véling.is . |
| Outcome: | The main purpose of the study is to examine how gender bias appears in English-Icelandic translations. |
Modelling Frequency, Attestation, and Corpus-Based Information with OntoLex-FrAC (2022.coling-1)
Copied to clipboard
| Challenge: | OntoLex-Lemon has become a de facto standard for lexical resources in the web of data. |
| Approach: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information. |
| Outcome: | This paper provides the first overall description of the emerging OntoLex module for Frequency, Attestations, and Corpus-Based Information (OntoLx-FrAC) it is intended to complement OntoLemon with the vocabulary to represent major types of information found in or automatically derived from corpora, for applications in both language technology and the language sciences. |
From text to talk: Harnessing conversational corpora for humane and diversity-aware language technology (2022.acl-long)
Copied to clipboard
| Challenge: | Informal social interaction is the primordial home of human language. |
| Approach: | They show that linguistically diverse conversational corpora can provide empirical foundations for flexible, localizable language technologies of the future. |
| Outcome: | The results suggest that even relatively small corpora can support robust generalizations about key aspects of interactional infrastructure. |
Sustainable Modular Debiasing of Language Models (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing debiasing methods modify all of the PLM parameters, which is costly and leads to (catastrophic) forgetting of useful language knowledge. |
| Approach: | They propose a modular debiasing approach based on dedicated adapters that inject adapter modules into the original PLM layers and update only the adapters. |
| Outcome: | The proposed approach is based on dedicated adapters and retains fairness even after large-scale training. |
Detection of Reading Absorption in User-Generated Book Reviews: Resources Creation and Evaluation (2020.lrec-1)
Copied to clipboard
Piroska Lendvai, Sándor Darányi, Christian Geng, Moniek Kuijpers, Oier Lopez de Lacalle, Jean-Christophe Mensonides, Simone Rebora, Uwe Reichel
| Challenge: | a new study aims to detect how and when readers are experiencing engagement with a literary work . empirical literary studies and language technology are used to investigate reading absorption . |
| Approach: | They annotated user-generated book reviews with reading absorption categories . they then performed supervised binary classification of the mental state of absorption . |
| Outcome: | The proposed corpus of user-generated reviews is compared with machine learning models and a benchmark corpus. |
What a Creole Wants, What a Creole Needs (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent efforts to improve the quality of high-resource languages focus on translating existing datasets into other languages, but this approach ignores that different language communities have different needs. |
| Approach: | They examine how things needed from language technology can change dramatically from one language to another. |
| Outcome: | The proposed method ignores that different language communities have different needs. |
A Tree Extension for CoNLL-RDF (2020.lrec-1)
Copied to clipboard
| Challenge: | CoNLL-RDF provides a bridge for popular oneword-per-line formats . main reasons for their popularity are the simplicity of tables and tab-separated values . |
| Approach: | They propose a technology that provides a bridge between knowledge graphs and natural language processing. |
| Outcome: | The proposed technology provides a bridge for popular one-word-per-line formats . it provides native support for word-level annotations, but not phrase structures or text structure . |
PARME: Parallel Corpora for Low-Resourced Middle Eastern Languages (2025.acl-long)
Copied to clipboard
Sina Ahmadi, Rico Sennrich, Erfan Karami, Ako Marani, Parviz Fekrazad, Gholamreza Akbarzadeh Baghban, Hanah Hadi, Semko Heidari, Mahîr Dogan, Pedram Asadi, Dashne Bashir, Mohammad Amin Ghodrati, Kourosh Amini, Zeynab Ashourinezhad, Mana Baladi, Farshid Ezzati, Alireza Ghasemifar, Daryoush Hosseinpour, Behrooz Abbaszadeh, Amin Hassanpour, Bahaddin Jalal Hamaamin, Saya Kamal Hama, Ardeshir Mousavi, Sarko Nazir Hussein, Isar Nejadgholi, Mehmet Ölmez, Horam Osmanpour, Rashid Roshan Ramezani, Aryan Sediq Aziz, Ali Salehi, Mohammadreza Yadegari, Kewyar Yadegari, Sedighe Zamani Roodsari
| Challenge: | UNESCO has identified 60 varieties of Middle Eastern languages as underrepresented . a limited availability of language technology perpetuates a cycle of digital exclusion . |
| Approach: | They develop a parallel corpora for eight severely under-resourced varieties in the region . they evaluate machine translation capabilities through zero-shot approaches and fine-tuning experiments . |
| Outcome: | The proposed model aims to improve the processing of the eight under-resourced languages in the Middle East. |